Papers with vision-language and multimodal tasks

1 papers
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods to learn biases from the perspective of model components are limited by their complexity and performance.
Approach: They propose a framework that incorporates causal mediation analysis to measure and map the pathways of bias generation and propagation within vision-language and multimodal tasks.
Outcome: The proposed framework is applicable to a wide range of vision-language and multimodal tasks and reduces bias by 22.03% and 9.04% in the MSCOCO and PASCAL-SENTENCE datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations